Papers by Josef Van Genabith
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages (2025.acl-srw)
Copied to clipboard
| Challenge: | Low-resource languages (LRLs) face significant challenges in natural language processing due to limited data. |
| Approach: | They evaluate adapter-based methods for adapting mLMs to low-resource languages . they use unstructured text and structured knowledge from ConceptNet to evaluate adapters . |
| Outcome: | The proposed methods outperform large language models and LLaMA-3 and deepSeek-R1 models on low training data. |
When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation (2026.eacl-short)
Copied to clipboard
| Challenge: | Large language models (LLMs) can be benchmark-contaminated, resulting in inflated scores that mask memorization as generalization. |
| Approach: | They use the FLORES-200 translation benchmark as a diagnostic to investigate cross-direction data contamination. |
| Outcome: | The proposed model can be cross-directional, boosting performance in unseen translation directions due to target-side memorization. |
Continual Learning in Multilingual Sign Language Translation (2025.naacl-long)
Copied to clipboard
| Challenge: | Despite the low translation quality of sign language, many machine learning approaches are still in its infancy. |
| Approach: | They propose to use continual learning for mul- tilingual SLT to improve translation quality. |
| Outcome: | The proposed methods outperform baseline and fine-tuning approaches in sign language translation. |